About
Head of the Data Science Platform at the Imagine Institute since 2012, I support researchers and clinicians in making the most of their data to advance research on rare genetic diseases in children. My daily work combines tool design, data structuring and analysis, application development, and methodological support for the Institute's teams and laboratories.
The platform provides the infrastructure and expertise needed to turn health data (clinical, biological) into knowledge that serves both research and diagnosis. In particular, we develop data warehouse and artificial intelligence solutions that help to better understand these diseases and to speed up patient care.
Our role is also to bridge the needs of the teams with the possibilities offered by data and digital technology: identifying a cohort, making use of a dataset, imagining a new tool, or making an analysis more reliable. This expertise is designed as a shared service, for the benefit of the Institute's entire community.
That is why we welcome your requests: a project to structure, data to make the most of, a methodological question, or an idea for an application… please don't hesitate to come and discuss it with us. The platform is here to serve you, and it is together that we will get the most out of it.
Scientific project
My work is organized around four complementary areas, ranging from the structuring of hospital data to diagnostic support and the generation of research hypotheses.
Area 1. Multimodal integration and information retrieval. This first area aims to bring together and make searchable data of very different kinds: clinical reports, biological and genomic data, and biomedical literature. The Dr. Warehouse project, a health data warehouse centered on clinical text, forms its foundation. It makes it possible to search and cross-reference the information contained in patient records across the entire Institute. The goal is to bring hospital data and knowledge drawn from the literature into dialogue within a single research environment.
Area 2. Information extraction from hospital data. The second area develops methods to transform the free text of medical records into usable data: entity recognition (symptoms, diagnoses, treatments), normalization to medical terminologies, and phenotype detection. These natural language processing methods feed into all of my other work by making the raw clinical material more reliable and better structured.
Area 3. Reducing diagnostic delay through computational patient representation. This third area tackles head-on one of the major challenges of rare diseases: the delay before diagnosis. By building a computational representation of the patient, a synthetic profile based on all of their data, I seek to measure similarity between patients and bring clinically similar cases closer together. The AIDY project fits into this diagnostic-support approach, helping to identify earlier those patients likely to share the same, as yet uncharacterized, genetic disease.
Area 4. Exploration and hypothesis generation. The final area explores the tools that allow researchers and clinicians to navigate the data and bring out new avenues of research. The Dr Cloud project opens up these analytical capabilities to a wider community by offering an accessible exploration environment. The aim is to move from a passive consultation of data to a genuine exploratory approach that generates clinical and scientific hypotheses.
Together, these four areas form a continuum: structuring the data, extracting meaning from it, bringing patients closer together for better diagnosis, and equipping research to make discoveries.